Skip to content
  • 0 Votes
    1 Posts
    214 Views
    A
    <p>According to OpenAI's latest billing rules, models such as GPT-5.6 Sol use tiered pricing for extra-long context requests.</p><p>When the context of a single request exceeds 272K tokens, input, cached reads, and output may be billed at higher-tier prices. This rule comes from OpenAI's official pricing and does not represent a temporary surcharge or abnormal charge from AI-ROUTER.</p><p></p><p>Using GPT-5.6 Sol's standard pricing as an example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Item Within 272K Over 272K</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ━━━━━━━━━━ ━━━━━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━━━━</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Input $5 / 1M tokens $10 / 1M tokens</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ────────── ─────────────────── ─────────────────</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Cached read $0.50 / 1M tokens $1 / 1M tokens</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ────────── ─────────────────── ─────────────────</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Output $30 / 1M tokens $45 / 1M tokens</code></pre><p></p><p>The final cost is still calculated based on the specific group, account multiplier, and other billing settings.</p><p></p><p>An x2 indicator in usage records means that the request triggered long-context tiered billing. This indicator mainly means that input and cached reads entered a higher tier; it does not mean that all costs for the entire request are simply doubled. Output may use a different tier multiplier.</p><p></p><p><strong> ## How to Avoid Triggering Tiered Billing</strong></p><p>If you do not need a 1M-token context window, it is recommended that you adjust Codex's context limit back to 272K or below.</p><p></p><p>Open:</p><p> ~/.codex/config.toml</p><p></p><p>Change the 1M configuration from the original article:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model = "gpt-5.6-sol"</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_context_window = 1000000</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_auto_compact_token_limit = 900000</code></pre><p> </p><p>to:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model = "gpt-5.6-sol"</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_context_window = 272000</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_auto_compact_token_limit = 250000</code></pre><p> </p><p>Save the file, restart Codex, and start a new session.</p><p></p><p>If you only want this to apply temporarily to a single session, you can use:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>codex -m gpt-5.6-sol -c model_context_window=272000 -c model_auto_compact_token_limit=250000</code></pre><p> It is also recommended that you:</p><p> - Enable automatic compaction promptly;</p><p> - Split extra-long tasks across multiple sessions;</p><p> - Reduce the amount of tool output returned at once;</p><p> - Regularly summarize and clear older conversation content.</p><p></p><p>If you do need a 1M-token context window, you can continue using the original configuration, but usage beyond 272K will be billed according to OpenAI's higher-tier pricing.</p><p></p><p>For the original configuration instructions, see: <a target="_blank" rel="noopener noreferrer nofollow" href="https://ai-router.dev/blog/post-g-677aa3e043912838-how-to-enable-a-1m-token-context-window-in-codex-for-gpt-5-6-sol">How to Enable a 1M-Token Context Window in Codex for GPT-5.6 Sol </a></p><p></p><p>Thank you for your understanding and support.</p>
  • 0 Votes
    1 Posts
    3 Views
    A
    <p>According to OpenAI's latest billing rules, models such as GPT-5.6 Sol use tiered pricing for ultra-long context requests.</p><p>When the context of a single request exceeds 272K tokens, input, cached reads, and output may be billed at higher tier prices. This rule comes from OpenAI's official pricing and is not a temporary surcharge or abnormal charge from AI-ROUTER.</p><p></p><p>Using GPT-5.6 Sol's standard prices as an example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Item Within 272K Over 272K</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ━━━━━━━━━━ ━━━━━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━━━━</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Input $5 / 1M tokens $10 / 1M tokens</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ────────── ─────────────────── ─────────────────</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Cached reads $0.50 / 1M tokens $1 / 1M tokens</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ────────── ─────────────────── ─────────────────</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Output $30 / 1M tokens $45 / 1M tokens</code></pre><p></p><p>The final cost is still calculated based on the specific group, account multiplier, and other billing settings.</p><p></p><p>If an x2 flag appears in usage records, it means that the request triggered tiered billing for long contexts. This flag mainly indicates that input or cached reads entered a higher tier; it does not mean that all charges for the request were simply doubled. Output may use a different tier multiplier.</p><p></p><p><strong> ## How to Avoid Triggering Tiered Billing</strong></p><p>If you do not need a 1M-token context window, we recommend adjusting Codex's context limit back to within 272K.</p><p></p><p>Open:</p><p> ~/.codex/config.toml</p><p></p><p>Change the 1M configuration from the original article:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model = "gpt-5.6-sol"</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_context_window = 1000000</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_auto_compact_token_limit = 900000</code></pre><p> </p><p>To:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model = "gpt-5.6-sol"</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_context_window = 272000</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_auto_compact_token_limit = 250000</code></pre><p> </p><p>Save the changes, restart Codex, and start a new session.</p><p></p><p>If you only want this to apply temporarily to a single session, you can use:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>codex -m gpt-5.6-sol -c model_context_window=272000 -c model_auto_compact_token_limit=250000</code></pre><p> We also recommend:</p><p> - Enable automatic compaction promptly;</p><p> - Split ultra-long tasks across multiple sessions;</p><p> - Reduce the amount of tool output returned at once;</p><p> - Periodically summarize and clear older conversation content.</p><p></p><p>If you do need a 1M-token context window, you can continue using the original configuration, but usage beyond 272K will be billed according to OpenAI's higher tier pricing.</p><p></p><p>For the original configuration instructions, see: <a target="_blank" rel="noopener noreferrer nofollow" href="https://ai-router.dev/blog/post-g-677aa3e043912838-how-to-enable-a-1m-token-context-window-in-codex-for-gpt-5-6-sol">How to Enable a 1M-Token Context Window in Codex for GPT-5.6 Sol </a></p><p></p><p>Thank you for your understanding and support.</p>
  • 0 Votes
    1 Posts
    203 Views
    A
    <p>According to OpenAI's latest pricing rules, models such as GPT-5.6 Sol use tiered pricing for ultra-long-context requests.</p><p>When the context of a single request exceeds 272K tokens, input, cached reads, and output may be charged according to higher-tier prices. This rule comes from OpenAI's official pricing and is not a temporary markup or abnormal charge from AI-ROUTER.</p><p></p><p>Using GPT-5.6 Sol's standard pricing as an example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Item Within 272K Over 272K</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ━━━━━━━━━━ ━━━━━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━━━━</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Input $5 / 1M tokens $10 / 1M tokens</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ────────── ─────────────────── ─────────────────</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Cached read $0.50 / 1M tokens $1 / 1M tokens</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> ────────── ─────────────────── ─────────────────</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> Output $30 / 1M tokens $45 / 1M tokens</code></pre><p></p><p>The final cost is still calculated based on the specific group, account multiplier, and other billing settings.</p><p></p><p>An x2 flag in usage records indicates that the request triggered tiered pricing for long contexts. This flag primarily indicates that input and cached reads entered a higher tier; it does not mean that all charges for the request are simply doubled. Output may use a different tier multiplier.</p><p></p><p><strong> ## How to Avoid Triggering Tiered Pricing</strong></p><p>If you do not need a 1M-token context window, we recommend adjusting Codex's context limit back to 272K or below.</p><p></p><p>Open:</p><p> ~/.codex/config.toml</p><p></p><p>Change the 1M configuration from the original article:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model = "gpt-5.6-sol"</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_context_window = 1000000</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_auto_compact_token_limit = 900000</code></pre><p> </p><p>to:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model = "gpt-5.6-sol"</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_context_window = 272000</code></pre><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code> model_auto_compact_token_limit = 250000</code></pre><p> </p><p>Save the changes, restart Codex, and start a new session.</p><p></p><p>If you only want the change to apply temporarily to a single session, use:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>codex -m gpt-5.6-sol -c model_context_window=272000 -c model_auto_compact_token_limit=250000</code></pre><p> We also recommend:</p><p> - Enable automatic compaction promptly;</p><p> - Split ultra-long tasks across multiple sessions;</p><p> - Reduce the amount of tool output returned at once;</p><p> - Regularly summarize and clean up older conversation content.</p><p></p><p>If you genuinely need a 1M-token context window, you can continue using the original configuration, but usage beyond 272K will be charged according to OpenAI's higher-tier pricing.</p><p></p><p>For the original configuration instructions, see: <a target="_blank" rel="noopener noreferrer nofollow" href="https://ai-router.dev/blog/post-g-677aa3e043912838-how-to-enable-a-1m-token-context-window-in-codex-for-gpt-5-6-sol">How to Enable a 1M-Token Context Window in Codex for GPT-5.6 Sol </a></p><p></p><p>Thank you for your understanding and support.</p>